fix(tier3): propagate Harbor dual-arm task suffixes and normalize canonical case ID resolution - #161
kweinmeister wants to merge 7 commits into
Conversation
…ll standards Signed-off-by: Karl Weinmeister <kweinmeister@google.com>
…ution - Propagate arm suffixes to `[task] name` in staged native `task.toml` files. - Decouple arm-suffix and attempt-suffix stripping with a commutative pipeline. - Strip external repository and namespace prefixes before canonicalizing IDs. - Protect case IDs retaining `skillevaluator-` or ending in `-with`/`-without`. - Enforce fail-fast runtime string validation for `arm_suffix` across adapter APIs. - Parameterize test suite and expand test matrix to cover all suffix variants. - Document dual-arm evaluation fixes in CHANGELOG.md. Signed-off-by: Karl Weinmeister <kweinmeister@google.com>
# Conflicts: # CHANGELOG.md
Signed-off-by: Karl Weinmeister <kweinmeister@google.com> # Conflicts: # CHANGELOG.md
rng1995
left a comment
There was a problem hiding this comment.
@kweinmeister Thanks for the contribution. The current normalization introduces cross-case identity collisions, and the native name rewrite can corrupt valid TOML. Please address the three inline findings with regressions before approval.
Local validation: 418 focused runtime, case-ID, metrics, adapter, collector, and security-attribution tests passed; Ruff passed. Separate reproductions exposed the reported gaps. All 17 reported CI checks pass.
There is also a merge conflict with main in CHANGELOG.md. Could you please resolve it so the updated PR can complete verification and move toward merge?
| return | ||
| content = task_toml.read_text(encoding="utf-8") | ||
|
|
||
| pattern = r'(?ms)(\[task\]\s*?\n(?:(?!\[)[^\n]*\n)*?\s*name\s*=\s*)(["\'])(.*?)\2' |
There was a problem hiding this comment.
[P2] Update the native task name structurally
The regex does not preserve valid TOML syntax or reliably target [task].name. For example, name = """nvidia/case-1""" is accepted by Harbor but becomes name = "-with-skill""nvidia/case-1""", which raises TOMLDecodeError. A valid [task] # task identity header also misses this pattern; if [metadata].name appears earlier, the fallback modifies that field and leaves the task name unchanged. Parse the document and update only task.name, preserving unrelated fields, and add cases for triple-quoted strings and commented section headers.
There was a problem hiding this comment.
Replaced the regex substitution in _append_native_task_name_suffix with tomllib.loads parsing and a token-aware value-span scanner for [task].name. It handles single and triple-quoted strings, inline comments on [task] headers, multiline arrays, and preceding [metadata].name or inline tables, then validates the updated document with tomllib.loads before writing. Added parameterized tests for these TOML variants.
There was a problem hiding this comment.
Still reproducible at current head 18cac4a. Harbor accepts a native task.toml containing task = { name = "nvidia/case-1-with-skill" }, but _find_toml_table_key_value_span finds no scalar assignment, so staging silently leaves that task name unchanged. With adjacent task directories case-1 and case-1-with-skill, a real Harbor 0.13.2 Oracle/Docker dual-arm run completed 4/4 trials with no exceptions; collection then reported Missing scored attempts for cases: case-1-with-skill and Excess scored attempts for cases: case-1 in the with-skill arm and omitted its pass@k and lift. Please handle valid inline task tables (and escaped quoted task/name keys) or fail staging before launching ambiguous names, with a dual-arm regression for the resulting report.
There was a problem hiding this comment.
@kweinmeister Thank you for your contribution and the follow-up. The original triple-quoted strings and commented headers now work, but I independently reproduced the remaining inline-table and escaped-key cases at 18cac4a: staging silently leaves task.name unchanged. Please handle these valid TOML representations or reject them before launching, and cover the resulting dual-arm collection behavior. Keeping this thread open.
Signed-off-by: Narendran Raghavan <nraghavan@nvidia.com>
|
Resolved the merge conflict with Local verification: 48 case-ID/report dependency tests passed; diff checks and Fern validation passed (authenticated redirects check skipped). The three code-review findings remain open, so this is not an approval. |
Signed-off-by: Karl Weinmeister <kweinmeister@google.com> # Conflicts: # CHANGELOG.md
…tructurally Signed-off-by: Karl Weinmeister <kweinmeister@google.com>
|
Merged main and pushed updates for the review comments:
|
| return Path(task_path).name | ||
| meta_id, dir_name = _result_task_path_entry_id(task_id.get("path")) | ||
| if meta_id: | ||
| return meta_id |
There was a problem hiding this comment.
[P2] Align native expected IDs with the metadata ID used here
A native task can have directory physical-case and [metadata].entry_id = "authored-entry"; staging accepts and preserves both. This branch returns authored-entry, but the runner passes staged directory names as expected_case_ids (runner.py:2367,2594). In a real Harbor 0.13.2 Oracle/Docker single-arm run, physical-case completed its trial with reward 1.0 and no exceptions. Collection nevertheless failed with Unexpected scored cases: authored-entry and Missing scored attempts for cases: physical-case, leaving pass@k empty. Please align the runner's expected IDs, collected identity, and dataset snapshot for native tasks, and cover a custom-only task whose metadata ID differs from its directory.
There was a problem hiding this comment.
@kweinmeister Thank you for your contribution. Reproduced at 18cac4a: a task under physical-case with metadata ID authored-entry is collected under its metadata ID, while the runner expects the directory ID, producing both unexpected-case and missing-attempt errors. Please use a consistent mapping across staging, expected IDs, collection, and the dataset snapshot, with a custom-only regression. This remains unresolved.
rng1995
left a comment
There was a problem hiding this comment.
@kweinmeister Thank you for your contribution and revisions. I verified and resolved two earlier findings: authoritative-ID attribution and authored arm-suffix handling. The inline/escaped TOML representation and metadata-versus-directory identity findings remain reproducible; I followed up in their existing threads.
Local validation: 382 focused tests passed, four skipped; diff checks passed. CI has 16 successful checks; Tier 2 macOS failed before tests during dependency installation because PyPI returned HTTP 503 for setuptools. Please rerun that job after addressing the remaining findings. No merge conflicts are reported, but the two code blockers still prevent approval.
Summary
When running Tier 3 Harbor dual-arm (
with-skill/without-skill) evaluations, native stagedtask.tomlfiles retained their base[task] namewithout the arm suffix, and result collection could fail to correlate paired trials when Harbor produced external namespace prefixes (repo/,org__), attempt suffixes (__attempt-N,-attempt-N), or arm suffixes (-with-skill,-without-skill) in varying orders.This change:
task.toml: Updates_rewrite_task_tomlandcopy_native_tasks_with_skill_modeinsrc/skillevaluator/tier3/harbor/adapter.py(wired fromsrc/skillevaluator/tier3/harbor/runner.py) so[task] namein staged nativetask.tomlfiles includes-with-skill/-without-skillalongside the staged task directory name.src/skillevaluator/tier3/harbor/collector.py(_strip_attempt_suffix,_strip_arm_suffix,_canonical_case_id) to strip external repository/namespace prefixes and commutatively strip attempt and arm suffixes in any order while preserving legitimate case IDs (such as those inexpected_ids, retainingskillevaluator-, or ending in-with/-without).arm_suffixinputs across adapter APIs and adds parameterized unit and integration tests intests/test_tier3_public_runtime.py.Verification
make lintmake testmake buildRelease Impact
CHANGELOG.md